[Fix] Fit the browser front's state and questions into Laya's window - #9
chaimaerachdi wants to merge 3 commits into
Conversation
56eac3a to
36175eb
Compare
Laya reads a 512 to 1024 token window, but the browser front sends every model the same state it sends Jev: a JSON object per element row, the full page text, and ten actions of history. On a real page that state fills the window well before a single instruction token is spent (docs/benchmarks.md shows 18,785-23,654 input tokens for Jev on the Allrecipes run), which is why `--model laya` routinely raises MODEL_SERVICE_CONFIG_ERROR on the browser front today. laya_state() folds a browser-shaped state before every call to LayaModel._decide: page.text dropped (the choice heads already carry each candidate's own text; the free-form dump is for the chat model's DONE answer, which Laya never writes), each element row rendered as one short line instead of a JSON object, and the last three actions kept instead of ten. On by default; anything that isn't the browser front's shape passes through unchanged (the tool front already fits). LAYA_COMPACT_BROWSER_STATE=0 turns it off. 34 unit tests (tests/test_decision_models_laya.py) cover the compaction itself, its wiring into LayaModel, and the env-var opt-out, all against FakeLayaAgent (no torch/weights needed). Full suite run against main: identical 19 pre-existing failures before and after this change (missing `ty` binary and other env-only gaps in this sandbox, unrelated to decision_models/laya.py). Not done here, and worth flagging: this closes the "state is bigger than the window" gap, not the "is Laya's window, even filled, actually fast enough end to end on a real page" question. That needs the real convaiinnovations/laya checkpoint (uv sync --extra laya), a live browser run, and a real latency number. See the PR description for exact commands. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Laya fits a question's instruction and all its options into one head_max_len
budget (192 by default). The browser front's target options are JSON objects,
so a 23-element click head left each option about six tokens,
`12: {"element": "[`, and Laya never saw an element's name.
laya_browser_question rewrites each browser question when the state is folded:
the instruction becomes the goal and the operation (the agent's rules, 446
tokens, are dropped), and each target option becomes its element's label and
value. Docs and .env.example give the window a browser run needs
(LAYA_MAX_LEN=1536, LAYA_HEAD_MAX_LEN=1024) and the measured result: the stock
checkpoint still answers DONE at the first step on Google Flights.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
0c470d9 to
e513234
Compare
|
Thanks for this. The diagnosis is solid: the six-token options ( 1.
Laya is asked to choose an operation from a list of values. 2. Truncation collapses distinct options into the same text With 28 characters per option and the Result lists and calendars (where labels share a long prefix) are exactly the pages the docs measure. Does Laya read the criteria keys, or only the option text? If only the text, it can't tell these apart. Possible fixes: keep the tail of the label instead of the head when labels collide, or add the index or 3.
Non-blocking
What I ran (macOS 26.5, Python 3.14; CI uses 3.11/3.13): Requesting changes for item 1; items 2 and 3 are up to you. |
- text_value (a goal, no operation) was asked "Which operation comes next?".
The rewrite now keys on the head: operation, <op>_target, text_value
("Which value should be typed into the field?").
- Options cut to 28 characters could become identical (a result list, a
calendar); two options with the same text now keep their key in front.
- A blocked row kept only the letter B; it now keeps its overlay's name, the
link to the button that closes it.
- is_browser_state() names the check laya_state and _decide share, in place
of the object-identity test.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Thanks for the careful review. Fixed in ecefeea:
|
Why
--model layaon a browser agent could not work as the front sent it. Two things did not fit Laya's window:full page text and ten actions of history, about 7K tokens per decision on Google Flights and 18,785 to 23,654
on Allrecipes (
docs/benchmarks.md). Laya reads 512 to 1024 tokens.head_max_lenbudget (192by default). Past it, every option is cut to an equal share. The browser front's target options are JSON
objects, so a 23-element click head left each option six tokens,
12: {"element": "[: Laya never saw anelement's name. The instruction also carried the agent's rules, 446 tokens.
How
Both folds are in
s1a/decision_models/laya.pyand run inLayaModel._decideonly when the state has the browserfront's shape; the tool front and the rails pass through unchanged.
laya_state()folds the state:page.textdropped, one short line per element row, the last three actions.laya_browser_question()folds each question: the instruction becomes the goal and the operation, and each targetoption becomes its element's label and value,
Where from? = Zurich.What
LAYA_COMPACT_BROWSER_STATE(default on;0,falseornoturns both folds off).206 tokens for the operation head, 73 to 101 for TYPE_TEXT and PRESS_ENTER, 92 to 915 for CLICK (a calendar page
offers 66 days), 98 to 1,002 tokens of state. So:
LAYA_MAX_LEN=1536 LAYA_HEAD_MAX_LEN=1024.docs/decision-models.md,docs/configuration.md,.env.example,CHANGELOG.md.laya_state()to shared code for Julia-1 as well.Results: the request fits, the stock checkpoint still fails
Google Flights, Zurich to London, one adult, economy. Same Chrome, CPU only (no GPU).
1536/1024(2026-09-28)typed-decisionscheckpoint.on 4 of 23 questions (the operation head and the chosen operation's target head).
So the fit problem is fixed, but Laya does not know how to drive a web form. Its model card names email triage,
routing, guardrails and moderation as what its checkpoints are for. Browser use would need a checkpoint fine-tuned on
browser steps. I recorded Jev on 12 more Google Flights routes as training data (all DONE, 145 decisions), but
fine-tuning on my CPU was too slow to finish (about 30 minutes per pass over 278 examples). A GPU would make it
possible.
Verification
uv run ruff format --check . && uv run ruff check .: 137 files formatted, all checks passed.uv run ty check: 1 diagnostic, ins1a/decision_models/cua.py, not touched here; the same onmain.uv run pytest -q --ignore=tests/test_browser_policy.py: 16 failed, 404 passed, 41 skipped.mainon thismachine: 16 failed, 390 passed. The same 12 failing tests by name on both (Windows environment), none new.
tests/test_browser_policy.pyfails to collect onmaintoo (openjiuwen.harness.schema.decision_policymissing in the installed
openjiuwen).uv run pytest tests/test_decision_models_laya.py -q: 43 passed, 2 skipped (fake agent, no weights).scripts/smoke.sh:smoke: ok.CHANGELOG.mdand the docs say what the code does now.LAYA_MAX_LEN=1536 LAYA_HEAD_MAX_LEN=1024 uv run s1a run flights --model laya --timeout 900(withuv sync --extra laya), result above.🤖 Generated with Claude Code